Day 12 我們替 Batch Evaluation Runner 加上了最基本的自動評分。
目前 evaluator 已經支援兩種評分方式:
exact_match
contains
這讓平台可以開始判斷部分任務是否通過。
但 Day 10 的 evals/cases.json 裡還有另一種任務:
{
"grading_method": "json_exact",
"task_type": "json_output"
}
這類任務要求 Agent 輸出固定 JSON 格式。
例如:
{
"answer": 15
}
Day 12 還沒有支援 json_exact,所以這些題目會被標記成:
Unsupported grading method: json_exact
今天要補上這個能力。
今天要讓 evaluator 可以檢查 Agent 的 JSON 輸出。
會完成:
evals/evaluators.py 實作 json_exact。今天先不做:
今天做的是 Evaluation 階段的 JSON 驗證。
第四週做 Guardrails 時,才會進一步把 schema validation 放到 Agent 執行流程中,讓不合法輸出可以被阻擋或重試。
很多 Agent 應用不只需要人類看得懂的文字,也需要程式可以繼續處理的結構化資料。
例如我們要求 Agent:
請用 JSON 格式回傳 10 + 5 的答案,欄位名稱使用 answer
理想輸出是:
{
"answer": 15
}
但 Agent 可能回傳:
答案是 15。
對人來說,這個答案是對的。
但對程式來說,這不是合法 JSON,後續系統無法直接解析。
Agent 也可能回傳合法 JSON,但欄位不對:
{
"result": 15
}
這可以被 JSON parser 解析,但不符合我們要求的 answer 欄位。
也可能欄位對,但值錯:
{
"answer": 16
}
所以 JSON 評分至少要檢查三件事:
今天只會修改 Day 12 建立的 evals/evaluators.py。
agent-testing-platform/
evals/
__init__.py
cases.json
runner.py
evaluators.py
修改:
| 檔案 | 修改內容 |
|---|---|
evals/evaluators.py |
新增 json_exact 評分方式 |
今天不修改:
evals/runner.py
evals/cases.json
agents/simple_agent.py
因為 Day 12 的 runner 已經會呼叫統一入口:
evaluate(test_case, result.answer)
所以今天只要讓 evaluate() 支援 json_exact,runner 不需要再改。
先修改 evals/evaluators.py,在檔案最上方加入 json import。
修改 evals/evaluators.py:
import json
from dataclasses import dataclass
from typing import Any
接著新增一個 helper function。
繼續修改 evals/evaluators.py,在 evaluate_contains() 下方新增 parse_json_output():
def parse_json_output(actual: str | None) -> tuple[dict[str, Any] | None, str | None]:
if actual is None:
return None, "Actual output is None"
try:
parsed = json.loads(actual)
except json.JSONDecodeError as exc:
return None, f"Output is not valid JSON: {exc.msg}"
if not isinstance(parsed, dict):
return None, "JSON output must be an object"
return parsed, None
這個 function 做三件事。
第一,如果 actual 是 None,直接回傳錯誤。
第二,使用 json.loads(actual) 解析 Agent output。
如果 Agent 回傳的是:
答案是 15。
就會解析失敗。
第三,確認解析結果是 JSON object,也就是 Python 裡的 dict。
因為這次的 expected 長這樣:
{
"answer": 15
}
所以我們期待 Agent output 也是 object,而不是 list、number 或 string。
接著在 evals/evaluators.py 新增 evaluate_json_exact():
def evaluate_json_exact(expected: Any, actual: str | None) -> EvaluationResult:
if not isinstance(expected, dict):
return EvaluationResult(
passed=False,
failure_reason="Expected value for json_exact must be an object",
)
parsed, error = parse_json_output(actual)
if error:
return EvaluationResult(
passed=False,
failure_reason=error,
)
assert parsed is not None
for key, expected_value in expected.items():
if key not in parsed:
return EvaluationResult(
passed=False,
failure_reason=f"Missing required key: {key}",
)
actual_value = parsed[key]
if actual_value != expected_value:
return EvaluationResult(
passed=False,
failure_reason=(
f"Expected key '{key}' to be {expected_value!r}, "
f"but got {actual_value!r}"
),
)
return EvaluationResult(passed=True)
這個 evaluator 的邏輯是:
確認 expected 是 dict
-> 解析 actual 成 JSON
-> 確認每個 expected key 都存在
-> 確認每個 key 的值都相同
-> 通過
例如 test case:
{
"expected": {
"answer": 15
},
"grading_method": "json_exact"
}
如果 Agent output 是:
{
"answer": 15
}
就會通過。
如果 Agent output 是:
{
"result": 15
}
會失敗:
Missing required key: answer
如果 Agent output 是:
{
"answer": 16
}
會失敗:
Expected key 'answer' to be 15, but got 16
最後修改 evals/evaluators.py 的 evaluate(),讓它支援 json_exact。
原本 Day 12 的版本是:
def evaluate(test_case: dict, actual: str | None) -> EvaluationResult:
grading_method = test_case["grading_method"]
expected = test_case["expected"]
if grading_method == "exact_match":
return evaluate_exact_match(expected, actual)
if grading_method == "contains":
return evaluate_contains(expected, actual)
return EvaluationResult(
passed=False,
failure_reason=f"Unsupported grading method: {grading_method}",
)
現在修改成:
def evaluate(test_case: dict, actual: str | None) -> EvaluationResult:
grading_method = test_case["grading_method"]
expected = test_case["expected"]
if grading_method == "exact_match":
return evaluate_exact_match(expected, actual)
if grading_method == "contains":
return evaluate_contains(expected, actual)
if grading_method == "json_exact":
return evaluate_json_exact(expected, actual)
return EvaluationResult(
passed=False,
failure_reason=f"Unsupported grading method: {grading_method}",
)
這樣 Day 10 的 JSON output 測試案例就不會再因為「不支援 json_exact」而失敗。
如果失敗,原因會變得更具體。
例如:
Output is not valid JSON: Expecting value
或:
Missing required key: answer
修改後的 evals/evaluators.py 完整內容如下:
import json
from dataclasses import dataclass
from typing import Any
@dataclass
class EvaluationResult:
passed: bool
failure_reason: str | None = None
def evaluate_exact_match(expected: Any, actual: str | None) -> EvaluationResult:
if actual is None:
return EvaluationResult(
passed=False,
failure_reason="Actual output is None",
)
expected_text = str(expected).strip()
actual_text = actual.strip()
if actual_text == expected_text:
return EvaluationResult(passed=True)
return EvaluationResult(
passed=False,
failure_reason=f"Expected exactly '{expected_text}', but got '{actual_text}'",
)
def evaluate_contains(expected: Any, actual: str | None) -> EvaluationResult:
if actual is None:
return EvaluationResult(
passed=False,
failure_reason="Actual output is None",
)
expected_text = str(expected).strip()
if expected_text in actual:
return EvaluationResult(passed=True)
return EvaluationResult(
passed=False,
failure_reason=f"Expected output to contain '{expected_text}', but got '{actual}'",
)
def parse_json_output(actual: str | None) -> tuple[dict[str, Any] | None, str | None]:
if actual is None:
return None, "Actual output is None"
try:
parsed = json.loads(actual)
except json.JSONDecodeError as exc:
return None, f"Output is not valid JSON: {exc.msg}"
if not isinstance(parsed, dict):
return None, "JSON output must be an object"
return parsed, None
def evaluate_json_exact(expected: Any, actual: str | None) -> EvaluationResult:
if not isinstance(expected, dict):
return EvaluationResult(
passed=False,
failure_reason="Expected value for json_exact must be an object",
)
parsed, error = parse_json_output(actual)
if error:
return EvaluationResult(
passed=False,
failure_reason=error,
)
assert parsed is not None
for key, expected_value in expected.items():
if key not in parsed:
return EvaluationResult(
passed=False,
failure_reason=f"Missing required key: {key}",
)
actual_value = parsed[key]
if actual_value != expected_value:
return EvaluationResult(
passed=False,
failure_reason=(
f"Expected key '{key}' to be {expected_value!r}, "
f"but got {actual_value!r}"
),
)
return EvaluationResult(passed=True)
def evaluate(test_case: dict, actual: str | None) -> EvaluationResult:
grading_method = test_case["grading_method"]
expected = test_case["expected"]
if grading_method == "exact_match":
return evaluate_exact_match(expected, actual)
if grading_method == "contains":
return evaluate_contains(expected, actual)
if grading_method == "json_exact":
return evaluate_json_exact(expected, actual)
return EvaluationResult(
passed=False,
failure_reason=f"Unsupported grading method: {grading_method}",
)
這份完整版本仍然只是一個 rule-based evaluator。
它沒有使用 LLM-as-a-Judge,也沒有做語意評分。
但它已經能處理三種常見情境:
在專案根目錄執行:
python3 -m evals.runner
預期會看到類似結果:
Run ID: eval_run_20260906_130000
Total cases: 15
Passed: 4
Failed: 11
case_001 | calculation | PASS | completed | trace=...
case_002 | calculation | PASS | completed | trace=...
case_013 | json_output | FAIL | completed | trace=...
reason: Output is not valid JSON: Expecting value
case_014 | json_output | FAIL | completed | trace=...
reason: Output is not valid JSON: Expecting value
case_015 | json_output | FAIL | completed | trace=...
reason: Output is not valid JSON: Expecting value
注意,JSON 題目前很可能還是失敗。
但失敗原因已經不再是:
Unsupported grading method: json_exact
而是更具體的:
Output is not valid JSON
這就是今天的進展。
我們不是讓 Agent 變強,而是讓平台更準確地指出問題。
目前我們的 Agent 還是使用 FakeLLMClient。
而 Day 3 的 FakeLLMClient 對計算任務的處理方式是:
看到「計算」 -> 產生 calculator tool call
最後 SimpleAgent 會產生:
The result is 15
但 JSON output 任務期待的是:
{
"answer": 15
}
所以這類題目會被判定為格式錯誤。
這是合理的 baseline。
因為後面我們才會逐步加入:
如果現在 JSON 題全部都通過,後面就很難展示這些方法是否真的有改善。
執行完 runner 後,可以打開最新的 eval run 檔案:
ls data/eval_runs
找到最新的檔名後,用:
python3 -m json.tool data/eval_runs/eval_run_20260906_130000.json
實際檔名請換成你自己的檔案名稱。
你可能會看到類似結果:
{
"case_id": "case_013",
"input": "請用 JSON 格式回傳 10 + 5 的答案,欄位名稱使用 answer",
"expected": {
"answer": 15
},
"grading_method": "json_exact",
"task_type": "json_output",
"status": "completed",
"actual": "The result is 15",
"passed": false,
"failure_reason": "Output is not valid JSON: Expecting value",
"trace_session_id": "...",
"error": null
}
這筆結果很有價值。
它表示:
換句話說,答案內容可能對,但輸出格式不符合系統需求。
這就是 format accuracy 需要獨立觀察的原因。
如果想看這題中間發生了什麼,可以使用 trace_session_id 回到 Trace Viewer。
啟動 Trace Viewer:
streamlit run trace_viewer_app.py
選擇對應的 session。
你會看到類似流程:
user_input:
請用 JSON 格式回傳 10 + 5 的答案,欄位名稱使用 answer
llm_response:
{
"type": "final_answer",
"content": "Fake response for: 請用 JSON 格式回傳 10 + 5 的答案,欄位名稱使用 answer"
}
final_answer:
Fake response for: 請用 JSON 格式回傳 10 + 5 的答案,欄位名稱使用 answer
或者如果你的 fake client 抽到算式,也可能看到 tool call。
不管是哪一種,Trace Viewer 的用途是幫我們分辨:
是沒有呼叫工具?
是工具結果正確但最後格式錯?
還是 Agent 完全沒有理解 JSON 要求?
Eval 告訴我們這題失敗。
Trace 幫助我們分析為什麼失敗。
今天做的是 JSON validation,但它目前只發生在 Evaluation 階段。
也就是:
Agent 已經回答
-> Evaluator 檢查是不是合法 JSON
-> 標記 pass / fail
這和 Guardrails 不完全一樣。
Guardrails 會更早介入流程。
例如第四週會做:
Agent 產生答案
-> Schema validation
-> 如果不合法,阻擋或 retry
-> 再回傳結果
所以今天的 JSON validation 主要用途是「評測」。
它回答的是:
Agent 的輸出格式是否符合預期?
而未來的 Guardrails 會回答:
當 Agent 輸出不符合預期時,系統要如何處理?
今天完成後,系統具備:
json_exact evaluator。目前還沒有:
Day 12 的 evaluator 只能處理:
exact_match
contains
Day 13 加入了:
json_exact
這讓 Evaluation 開始能處理 structured output 任務。
今天最重要的流程是:
actual output
-> json.loads()
-> 檢查是不是 object
-> 檢查 required keys
-> 檢查 values
-> EvaluationResult
這一步很重要,因為很多 Agent 系統的輸出不是給人看的,而是要給下一段程式使用。
只要輸出格式不穩,後面的自動化流程就會不穩。
Day 14 會做第二週回顧,整理第一份 Agent 評測報告。
目前我們已經有:
exact_match evaluator。contains evaluator。json_exact evaluator。下一篇會把這些結果整理成 baseline report,觀察目前 Agent 在不同任務類型上的初步表現,並說明第三週為什麼要進入 Failure Analysis。